Back

JNCI: Journal of the National Cancer Institute

Oxford University Press (OUP)

Preprints posted in the last 90 days, ranked by how well they match JNCI: Journal of the National Cancer Institute's content profile, based on 19 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Genetic and Shared Environmental Influences on Cancer Risk and Cross-Cancer Associations in Nordic Twins

Harris, J. R.; Clemmensen, S. B.; Adami, H.-O.; Mucci, L. A.; Kaprio, J.; Hjelmborg, J. v. B.

2026-06-22 epidemiology 10.64898/2026.06.18.26355861 medRxiv
Top 0.1%
22.1%
Show abstract

The relative contributions of genetic and shared environmental influences to cancer risk and cross-cancer associations remain poorly understood. We analyzed data from 222,530 same-sex twins from Denmark, Finland, Norway, and Sweden in the Nordic Twin Study of Cancer, including 43,060 incident cancers over a median follow-up of 41.6 years. Using a target trial framework, biometric modeling, and competing-risk adjustment, we estimated familial risk, heritability, and shared environmental contributions across 35 cancer sites. Lifetime cancer risk was 36.5%, increasing to 51.4% in monozygotic (MZ) twins and 45.3% in dizygotic (DZ) twins with an affected co-twin. Overall cancer risk was explained by heritable (28%) and shared environmental (40%) influences. Heritability was highest for prostate (42%), non-melanoma skin (24%), and breast (18%) cancers. Cross-cancer analyses revealed extensive overlap in the genetic and shared environmental factors across sites, consistent with widespread pleiotropy and shared environmental susceptibility. Prostate cancer exhibited the strongest genetic overlap with rectum/anus (12%) and kidney (11%) cancers, whereas co-shared environmental influences were most pronounced for breast-lung (11%), prostate-bladder (11%), and prostate-lung (12%) cancers. These findings show pervasive genetic overlap across cancers at different sites and emphasize the importance of incorporating familial shared environmental exposures into cancer risk prediction and prevention strategies.

2
Race and Socioeconomic Status Impact Survival from Early and Late-Onset Colorectal Cancer

Purrington, K.; Hsieh, M.-C.; Patil, S.; Mabvakure, B.; Ahn, J.; Zhang, R.; Ruterbusch, J. J.; Samdani, R.; Lee, G.; Wenzlaff, A.; Latif, S.; Dash, C.; Sartor, M.; Schwartz, A. G.; Stoffel, E. M.; Rozek, L. S.

2026-06-29 epidemiology 10.64898/2026.06.24.26356439 medRxiv
Top 0.1%
19.5%
Show abstract

Background: Colorectal cancer (CRC) disproportionately affects non-Hispanic Black (NHB) Americans compared to Non-Hispanic White (NHW), with more cases arising before age 50. Racial disparities in outcomes reflect complex interactions among healthcare access, socioeconomic factors, and structural racism, yet analyses linking individual-level data for these factors to survival remain limited. Methods: We examined overall and CRC-specific survival among NHB and NHW patients diagnosed between 2013 and 2022 enrolled in the Disparities and Cancer Epidemiology (DANCE) cohort, a population-based study of CRC in metropolitan Detroit and Louisiana. Multivariable Cox regression and competing-risks models were used to assess the roles of race, age of onset, neighborhood deprivation, and stage on survival outcomes. Results: Among 1,019 CRC cases (57% NHB, 43% NHW), NHB patients were more likely to reside in high-deprivation neighborhoods, report lower household incomes, and present with right-sided tumors, though stage at diagnosis did not differ by race. In multivariable analysis, stage was the strongest predictor of survival, while neighborhood deprivation (per 10-unit ADI increase: HR = 1.14) was independently associated with worse survival; NHB race was not significantly associated with survival after adjustment. Younger age at diagnosis was associated with a survival advantage in regional-stage disease but paradoxically with worse survival in distant-stage disease, and higher deprivation predicted worse survival in both local and distant but not regional stage. Conclusion: Our study shows that socioeconomic factors, as measured by ADI and household income, accounts for some, but not all, of the disparities in survival between NHB and NHW CRC cases.

3
Integration of lung tissue proteomics and genome-wide association data to identify lung cancer susceptibility proteins and potential drug targets

Xu, S.; Shi, J.; Shu, X.-O.; Tao, R.; Dou, Y.; Guo, X.; Wen, W.; Yang, Y.; Zhang, B.; Wu, J.; Deppen, S. A.; Li, B.; Zheng, W.; Long, J.; Cai, Q.

2026-06-22 epidemiology 10.64898/2026.06.18.26355973 medRxiv
Top 0.1%
15.1%
Show abstract

Background: Proteins directly impact disease development and act as drug targets. Therefore, we integrated genomic and lung tissue proteomics data to identify lung cancer susceptibility proteins, elucidating genetic mechanisms and candidate drug targets. Method: We profiled the proteome and genome in non-neoplastic lung tissue from 200 lung cancer patients. Using this data, we constructed genetic models to predict abundance across the proteome in lung tissue. We applied these models to genome-wide association study (GWAS) data from 55,174 lung cancer cases and 1,294,174 controls to evaluate their associations with the risk of lung cancer, overall and by major histological subtypes. Bayesian colocalization and Mendelian randomization (MR) analyses were used to prioritize putative causal proteins, which were cross-referenced with three main drug-protein databases to identify potential therapeutic targets. Results: We identified 29 proteins associated with lung cancer risk at a false discovery rate < 5%, including 25 for overall lung cancer, two (AQP3 and IL18) specifically for adenocarcinoma, and another two (HMGN2 and HLA-DMB) for squamous cell carcinoma. Of them, genes encoding 17 proteins reside at least 2Mb away from any known GWAS risk loci, including 14 for overall lung cancer (HYI, GPX1, GMPPB, DSP, HDDC2, MTCH2, SUOX, JMJD7, PDIA3, IL16, IQGAP1, SULT1A2, ARHGAP27, and TYMP) and three for subtypes (AQP3, IL18, and HMGN2). Among the 12 proteins located within the known risk loci, EPHX2, CLDN18, PSMD5, and CYP2S1 proteins showed an association independent of the proximal GWAS-identified lead variant. Colocalization and/or MR analysis suggested 11 potential causal proteins. Five of these candidate causal proteins (DSP, CLDN18, IQGAP1, IL18 and TYMP) are targeted by nine drugs already approved by the FDA or in phase III trials. Conclusion: Our study identified novel lung cancer susceptibility proteins and potential drug targets, offering valuable insights into lung cancer biology and future translational utilities.

4
Performance of family history-based colorectal cancer screening criteria by race and age at diagnosis in the Disparities and Cancer Epidemiology (DANCE) study

Purrington, K.; Martin, C.; Wenzlaff, A. S.; Ruterbusch, J. J.; Patil, S.; Pandolfi, S. S.; Samayoa, I.; Schwartz, A. G.; Hsieh, M.-C.; Stoffel, E. M.; Rozek, L. S.

2026-06-19 epidemiology 10.64898/2026.06.16.26355827 medRxiv
Top 0.1%
11.3%
Show abstract

Importance: Family history (FH) and age are the primary criteria employed for early colorectal cancer (CRC) risk stratification. We evaluated how well these criteria identify individuals diagnosed with CRC across age and racial groups. Objective: To evaluate the performance of FH and age based screening criteria for identifying individuals with CRC, with attention to differences by race and age at diagnosis. Design, Setting, and Participants: This case control and case only analysis used data from the Disparities and Cancer Epidemiology (DANCE) cohort, a population based study of invasive CRC cases diagnosed from 2013 to 2022, recruited through the Metropolitan Detroit Cancer Surveillance System and the Louisiana Tumor Registry. Analyses included 1,158 non-Hispanic Black (NHB) and non-Hispanic White (NHW) CRC cases and 1,434 cancer-free controls from the Inflammation Health and Lung Epidemiology (INHALE) study, enrolled from the same Detroit catchment area. Data were analyzed in 2025. Exposures: Self reported cancer FH among first-degree (FD) relatives and grandparents, summarized into three FH-based screening criteria: at least one FD relative with CRC (colon early-screening criterion), any FH of Lynch syndrome related cancers, and meeting NCCN criteria for Lynch syndrome genetic testing. Main Outcomes and Measures: Proportion of cases meeting each FH based screening criterion stratified by race and age at diagnosis (<45, 45 - 49, 50 - 64, and <65 years); case only odds ratios for younger age at diagnosis; and case control odds ratios for CRC associated with each criterion, with race-by-age interaction tested. Results: Cancer FH burden differed by age at diagnosis across both racial groups. First degree (FD) CRC FH was highest among NHB CRC cases diagnosed before age 45 (22.6%) and lowest in those diagnosed at ages 45-49 (8.2%), while NHW participants reported more CRC FH with older age at diagnosis (p interaction=0.011). In case control analyses, having at least one FD relative with CRC was associated with higher odds of CRC before age 45 among NHB (OR=1.44, 95% CI 1.09 - 1.89) but not NHW individuals. The proportion of cases diagnosed before age 45 with a FD CRC FH was low, though markedly higher in NHB than NHW individuals (22.6% vs. 4.0%). While the proportion was slightly higher when including FH of any Lynch syndrome-related cancers (NHB: 24.5%, NHW: 10.0%), the proportion of controls with a FD FH of these cancers also increased. Conclusions and Relevance: Current family history-based criteria fail to identify the majority of individuals diagnosed with CRC before age 45, with performance varying substantially by race, highlighting the urgent need for more equitable and effective approaches to early-onset CRC risk stratification.

5
Optimal LDCT screening for never-smoking Asian women using integrated polygenic and environmental risk: a microsimulation modelling study

Kowada, A.

2026-08-19 oncology 10.64898/2026.08.18.26360665 medRxiv
Top 0.1%
8.0%
Show abstract

Objective To identify optimal initiation ages and screening intervals for low-dose computed tomography (LDCT) screening among never-smoking Asian women using an integrated polygenic risk score (PRS)-environmental tobacco smoke (ETS) risk model, and to evaluate the cost-effectiveness of alternative screening strategies at these optimized ages. Design Integrated PRS-ETS microsimulation modelling. Setting Japan. Participants Never-smoking women stratified into eight risk groups defined by combinations of PRS levels and ETS exposure. Interventions LDCT screening at intervals of 1 to 10 years, annual chest radiography (CXR), or no screening. Main outcome measures Costs, quality-adjusted life years (QALYs), incremental cost-effectiveness ratios (ICERs), net monetary benefits, lung adenocarcinoma incidence and mortality, and optimal LDCT initiation ages. Sensitivity analyses used a willingness-to-pay threshold of US$50,000 per QALY gained. Results Optimal initiation ages ranged from 40 to 55 years across the eight PRS-ETS risk groups, with higher PRS-ETS risk associated with younger optimal initiation ages. Annual LDCT was the most cost-effective strategy across all PRS-ETS risk strata, yielding an ICER of US$40,471 per QALY in the lowest risk stratum and becoming cost-saving in higher risk strata. Over a lifetime, annual LDCT averted 8,534 lung adenocarcinoma deaths compared with annual CXR and 14,940 deaths compared with no screening. Conclusions Tailoring LDCT initiation age across integrated PRS-ETS risk groups maximizes mortality reduction achievable with cost-effective annual LDCT screening among never-smoking Asian women. These findings highlight an urgent limitation of global lung cancer screening guidelines that rely exclusively on smoking history and provide policy-ready evidence supporting the integration of PRS and ETS into future recommendations for precision LDCT screening for never-smoking populations.

6
Cannabis use and Cancer: Dissecting genetic causality for site-specific risks through two-sample Mendelian Randomization

Lukhere, E.; Kachingwe, B.; Kipandula, W.; Chiphangwi, N.; Singini, M. G.; Kamiza, A. B.

2026-08-13 genetic and genomic medicine 10.64898/2026.08.12.26360176 medRxiv
Top 0.1%
7.9%
Show abstract

Background: The prevalence of cannabis use is increasing at an alarming rate owing to its legalization and decriminalization in some countries. Epidemiological evidence on the association between cannabis use and cancer is inconsistent and conflicting. Herein, we performed two-sample Mendelian randomization (MR) to investigate whether cannabis use is causally associated with site-specific cancers in individuals of European ancestry. Methods: We identified 22 independent genetic variants strongly associated with cannabis use (p-value < 5 x 10-8) in a large meta-analysis of genome-wide association studies of individuals of European ancestry. Genome-wide association summary-level data on site-specific cancers were obtained from individuals of European ancestry in FinnGen, Finland. MR analyses were performed using the inverse-variance weighted (IVW) and multivariable method. Sensitivity analyses were performed using the simple median, weighted median, MR-Egger, and MR pleiotropy residual sum and outlier methods. Results: Our multivariable IVW analyses adjusted for cigarette smoking found that genetic liability to cannabis use was causally associated with esophageal cancer (odds ratio [OR] =1.74, 95% confidence interval [CI] =1.29-2.15, p-value =0.013) and lung cancer (OR=1.35, 95% CI = 1.13-1.58, p-value =0.009). However, genetic liability to cannabis use exerted a protective effect against pancreatic cancer (OR=0.77, 95% CI =0.57-0.91, pvalue=0.032) in individuals of European ancestry in the FinnGen. Our sensitivity analyses found no evidence of horizontal pleiotropy between cannabis use and site-specific cancers. Conclusion: We found that genetic liability to cannabis use was associated with esophageal, lung, and pancreatic cancers in individuals of European ancestry.

7
The urinary-metabolite-based lung cancer index (uLCI): an interpretable machine-learning risk model for early-stage disease

Khan, M. A.; Mathe, E. A.; Pine, S. R.; Gonzalez, F. J.; Harris, C. C.; Wang, X. W.; Patel, D. P.

2026-06-29 oncology 10.64898/2026.06.26.26356700 medRxiv
Top 0.1%
6.7%
Show abstract

Background Five-year survival from lung cancer exceeds 60% at stage I-II but falls below 10% once metastasis occurs. Low-dose CT (LDCT) screening reduces mortality in heavy smokers but carries a false-positive rate of approximately 29% and is restricted to smoking-based eligibility, leaving most cases undetected. We aimed to develop and independently validate an interpretable machine-learning urinary metabolite risk index (uLCI) for non-invasive lung cancer detection. Methods Four urinary metabolites - creatine riboside (CR), N-acetylneuraminic acid (NANA), 27-nor-5-beta-cholestane-3-alpha,7-alpha,12-alpha,24,25-pentol (CP), and cortisol sulfate (CS) - and three clinical variables (age, race, smoking) were integrated by Lasso-regularised logistic regression into a uLCI score. The model was developed under 10-fold cross-validation in the NCI-Maryland (NCI-MD) cohort (n=845; 470 controls, 375 cases, stages I-IV) and applied without refitting to the independent Colorado Lung Cancer Cohort (n=488; 211 controls, 277 cases). Analyses were prespecified; reporting followed TRIPOD+AI. Findings uLCI achieved an area under the curve (AUC) of 0.906 (95% CI 0.887-0.926) in NCI-MD and 0.748 (0.701-0.793) in the independent Colorado cohort. Scores rose monotonically across stages in both cohorts (Spearman rho=0.69 and 0.45; both p<0.0001). Stage-specific discrimination was preserved from stage I to IV (NCI-MD 0.900-0.927; Colorado 0.722-0.843). Net reclassification improvement over clinical variables was 1.24 (1.14-1.36) and 0.74 (0.56-0.90). uLCI tertiles stratified post-resection survival in stage I-II disease (adjusted hazard ratio 2.03, 1.26-3.27). Interpretation uLCI is an independently validated, interpretable urinary risk index that detects lung cancer across all stages, with monotonic stage progression and post-resection prognostic value. Its false-positive rate compares favourably with published estimates for LDCT and cell-free-DNA assays, supporting prospective head-to-head evaluation as a non-invasive triage tool, including in screening-ineligible populations. Funding Intramural Research Program, Center for Cancer Research, National Cancer Institute, US National Institutes of Health.

8
Associations between occupations and the occurrence of sarcomas: results of the French population-based case-control study ETIOSARC

GRAMOND, C.; Guillemin, L.; Coureau, G.; Systchenko, T.; Hammas, K.; GASH Illescas, A.; Delafosse, P.; Blay, J.-Y.; Ducimetiere, F.; Penel, N.; Toulmonde, M.; Le Loarer, F.; de Pinieux, G.; Baldi, I.; Monnereau, A.; Lacourt, A.; Mathoulin-Pelissier, S.; Amadeo, B.

2026-07-22 epidemiology 10.64898/2026.07.20.26358458 medRxiv
Top 0.1%
5.7%
Show abstract

Objective Sarcomas are rare tumors of connective tissue that can develop in soft-tissue, viscera organs, or bones. Previous occupational studies have mainly focused on men and soft-tissue sarcomas. This study describes associations between occupations and sarcomas, in men and women, for soft-tissue sarcomas (STS), including visceral sarcomas, and bone sarcomas (BS). Methods The ETIOSARC study was a multicenter case-control study conducted across six French geographical areas between 2019 and 2023. Occupational histories were collected by interview and coded according to the International Standard Classification of Occupations (2008). Conditional logistic regression models were applied to estimate odds ratios with 90% confidence intervals for each occupation included in the study. Results A total of 374 male cases (336 STS and 38 BS), 371 female cases (335 STS and 36 BS), and 1,387 controls were included. Positive associations were observed among male STS for painters (OR=5.37, 90% CI=1.83-15.71), waiters (OR=4.54, 90% CI=1.88-10.97), and among male BS for electricians (OR=3.67, 90% CI=0.91-14.86). Among women, increased STS risk was found among real estate agents (OR=3.48, 90% CI=1.03-11.79), food preparation assistants (OR=3.47, 90% CI=1.09-11.04), and administrative and executive secretaries (OR=2.37, 90% CI=1.45-3.86). For BS, a higher risk was observed among sales workers (OR=5.61, 90% CI=1.65-18.99). Conclusions Our study highlights hitherto unreported occupational associations with sarcomas, among both men and women, as well as the importance of sex-stratified analyses. Further research should aim to confirm these results and disentangle the respective roles of occupational exposures, environmental factors, and lifestyle characteristics.

9
Loss to Follow-Up Among Patients with Kaposi Sarcoma at the Ocean Road Cancer Institute, Tanzania: A Fine-Gray Competing-Risks Analysis of Death as a Competing Event

Lugina, E. L.; Mwita, C. J.; Nyamhanga, T. L.; Lidenge, S. J.; Ngowi, J. R.; Kahesa, C. L.; Wood, C.; Mwaiselage, J. D.

2026-08-28 epidemiology 10.64898/2026.08.26.26361391 medRxiv
Top 0.1%
5.6%
Show abstract

Purpose Kaposi sarcoma (KS) remains one of the most common HIV-associated malignancies in sub-Saharan Africa (SSA). While loss to follow-up (LTFU) has been well documented among KS patients managed within HIV primary care, little is known about retention after patients transition into specialized oncology care, where treatment pathways, toxicities, costs, and follow-up schedules differ substantially. Because LTFU is unlikely to occur at random, patients who disengage from care may differ systematically from those retained with respect to disease severity, treatment response, and mortality risk, potentially biasing survival estimates and underestimating cancer-related mortality. This study aimed to estimate the cumulative incidence of LTFU among patients with KS receiving care at Tanzanias national cancer referral center, accounting for death as a competing event, and to identify factors associated with LTFU. Methods This retrospective cohort study included 251 patients with KS treated at Ocean Road Cancer Institute (ORCI) between January 2021 and December 2023. The primary outcome was LTFU, with death treated as a competing event. Cumulative incidence of LTFU at 6, 12, 18, and 24 months was estimated using the cumulative incidence function. Predictors of LTFU were assessed using univariable and multivariable Fine-Gray subdistribution hazards regression. Results Among 251 patients, 214 (85.3%) had epidemic (HIV-associated) KS and 37 (14.7%) had endemic (non-HIV-associated) KS. Males accounted for 62.2%. Accounting for death as a competing event, the cumulative incidence of LTFU was 27.6% (95% CI, 22.0-33.1) at 6 months, 36.4% (95% CI, 30.5-42.4) at 12 months, 43.3% (95% CI, 37.2-49.4) at 18 months, and 46.6% (95% CI, 40.4-52.7) at 24 months. In multivariable Fine-Gray regression, absence of oral involvement (adjusted subdistribution hazard ratio [aSHR], 0.37), reachable telephone contact (aSHR, 0.54), and initial chemotherapy rather than radiotherapy (aSHR, 0.44) were independently associated with lower risk of LTFU. Conclusion Nearly half of patients with KS were LTFU within two years, substantially limiting reliable assessment of cancer outcomes in this setting. Strengthening retention strategies, including maintaining reliable patient contact information and implementing routine phone-based follow-up, may offer feasible, scalable approaches to improve continuity of care, enhance survival monitoring, and strengthen cancer surveillance in resource-limited settings.

10
The Role of Distress-related Metabolic Dysfunction in Ovarian Cancer Development: a pooled case-control study

Lin, N.; Balasubramanian, R.; Menichetti, G.; Eliassen, H.; Trabert, B.; Avila-Pacheco, J.; Townsend, M. K.; Terry, K. L.; Clish, C. B.; Tworoger, S. S.; Zeleznik, O. A.

2026-08-31 epidemiology 10.64898/2026.08.27.26361473 medRxiv
Top 0.1%
5.6%
Show abstract

Background: Evidence suggests chronic distress influences ovarian cancer (OC) etiology and metabolomic profiles. Here, we evaluated the association of a metabolite-based distress score (MDS) and OC risk. Methods: We included two matched case-control studies nested within the Nurses' Health Studies (N=584) and the Prostate, Lung, Colorectal, and Ovarian Cancer Screening Trial (N=348). Metabolites were measured 3-27 years before diagnosis using liquid-chromatography tandem mass spectrometry. We examined the association of quintiles of MDS and 19 constituent metabolites with OC risk using unconditional logistic regression and stratified by tumor histotype, menopausal status, and age at diagnosis. Results: We observed women in the highest versus lowest quintile of MDS had an increased OC risk (OR=1.62,95%CI=1.03-2.54,ptrend=0.07), and type 2 tumors (OR=1.71,95%CI=1.03-2.83,ptrend=0.11). Associations were suggestively stronger for premenopausal and <69-year-old women, and driven by pseudouridine, and N2,N2-dimethylguanosine. Conclusion: Our findings suggest chronic distress-associated metabolic dysregulation may represent a novel OC risk factor, especially among younger women.

11
Biological processes linking soft drink consumption with site-specific cancer risk within the Global Cancer Update Programme (CUP Global)

Fontvieille, E.; Ahmadi, N.; Mahamat-saleh, Y.; Hashem, N.; Lauby-Secretan, B.; Gunter, M. J.; Tabung, F. K.; Turner, S. D.; Kok, D. E.; Jones, L.; Herceg, Z.; Simpson, R. J.; Chan, D.; Tsilidis, K. K.; Jayedi, A.; Clary, C.; Croker, H.; Mitrou, P.; Riboli, E.; Hursting, S.; Lewis, S. J.; Dossus, L.

2026-07-16 epidemiology 10.64898/2026.07.13.26356051 medRxiv
Top 0.1%
5.6%
Show abstract

This review evaluates the biological pathways linking soft drink consumption with the risk of several cancers within the framework of the Global Cancer Update Programme (CUP Global). Soft drink consumption has been associated with increased risk of multiple cancers, and glucose or insulin dysregulation has been proposed as a potential underlying mechanism. We applied a three-stage framework. In the first stage, we identified insulin sensitivity as the key biological process potentially linking soft drink consumption (sugar-sweetened or artificially sweetened) to cancer risk, with glucose-related and insulin-related biomarkers as potential intermediate phenotypes, using a combination of expert knowledge and a web-based text mining tool. In the second stage, we conducted targeted PubMed searches to identify studies examining associations between consumption of soft drinks and these intermediate phenotypes (IPs) and between these IPs and the risk of several cancers in adult humans. In the third stage, the evidence was evaluated by the Expert Committee on Cancer Mechanisms (MEC), who assessed the strength of the evidence for these associations. The MEC concluded that there was weak evidence supporting a role of glucose or insulin-related processes as a potential mechanistic pathway linking the consumption of sugar-sweetened or artificially sweetened beverages to the risk of various cancers evaluated.

12
A Non-Invasive Urinary Bile-Acid Marker for Never-Smoker Lung Cancer

Chung, S.; Liu, H.; Khan, M.; Patel, T. S.; Blackman, B.; Swenson, R. E.; Pine, S. R.; Gonzalez, F. J.; Harris, C. C.; Patel, D. P.

2026-07-31 oncology 10.64898/2026.07.27.26359037 medRxiv
Top 0.1%
5.5%
Show abstract

Introduction: Lung cancer in never-smokers is a growing, biologically distinct entity lacking non-invasive markers. Established urinary markers - creatine riboside (CR) and N-acetylneuraminic acid (NANA) - report tumor-intrinsic metabolism, not carcinogen processing. We investigated 27-nor-5{beta}-cholestane-3,7,12,24R,25S-pentol glucuronide (CPG), a bile-acid glucuronide linked to aryl-hydrocarbon-receptor (AhR)/CYP xenobiotic metabolism. Methods: Urinary CPG was quantified by UPLC-tandem mass spectrometry in an exploratory (NCI-Maryland; n=846) and validation (Colorado; n=505) cohort of non-small-cell lung cancer cases and frequency-matched controls. Associations with case status, smoking stratum, survival, and discrimination were assessed, using tumor RNA sequencing (n=83) and gene-set enrichment analysis (GSEA). Results: Urinary CPG was higher in cases than controls in both cohorts (P<0.0001). In never-smokers, cases exceeded smoking-matched controls (P<0.001 and P<0.0001), indicating elevation independent of tobacco exposure. After mutual adjustment for CR and NANA, CPG remained independently associated with case status (exploratory OR 1.58, 95% CI 1.15-2.16; validation OR 3.92, 95% CI 2.47-6.29), with a modest gain in discrimination. High CPG identified never-smokers with worse survival in both cohorts (P<0.001 and P=0.04), remaining significant after multivariable adjustment only in the exploratory cohort. GSEA showed AhR/CYP xenobiotic and Nrf2 oxidative-stress enrichment in high-CPG tumors; the CPG aglycone carried disease-specific 24R,25S stereochemistry. Conclusions: Urinary CPG was associated with NSCLC in two retrospective case-control cohorts, including in a smoking-matched never-smoker comparison. High CPG also identified never-smokers with worse survival, remaining independently prognostic after adjustment in the exploratory cohort. Tumor expression does not establish tissue of origin. Prospective validation against CR and NANA is required.

13
Performance of general-population breast cancer risk prediction models in an international consortium

Brantley, K. D.; Ahearn, T. U.; Norton, E. L.; MacInnis, R.; Palmer, J. R.; Fortner, R. T.; Vachon, C. M.; Beane-Freeman, L.; Berrington de Gonzalez, A.; Frost, R.; Bertrand, K. A.; Zirpoli, G.; Neuhouser, M. L.; Barnett, M.; Teras, L. R.; Hodge, J. M.; Patel, A. V.; Bodelon, C.; Lacey, J. V.; Spielfogel, E. S.; Rohan, T. E.; Kirsh, V. A.; Langseth, H.; Tsuruda, K. M.; Milne, R. L.; Haiman, C.; Scott, C. G.; Eliassen, A. H.; Rosner, B.; Willett, W. C.; Romanos-Nanclares, A.; Chen, Y.; Wu, F.; Zheng, W.; Long, J.; O'Brien, K. M.; Sandler, D. P.; Kitahara, C. M.; Linet, M. S.; Anderson, G.; Lars

2026-08-23 epidemiology 10.64898/2026.08.20.26360899 medRxiv
Top 0.1%
5.3%
Show abstract

Background: Several breast cancer (BC) risk prediction models have been developed to provide personal risk assessments. Though individually validated, their performance has not been systematically evaluated across a wide range of populations or ages. Methods: We harmonized individual-level baseline questionnaire data and incident BC diagnoses from 21 cohorts from North America, Europe, and Australia participating in the Breast Cancer Risk Prediction Project. Five-year absolute risk of invasive BC was estimated for five established risk prediction models using classical risk factors only. Discrimination was evaluated by area under the curve (AUC). Calibration was assessed using average and risk-decile specific expected to observed (E/O) ratios. Performance metrics were meta-analyzed across cohorts and models. Metaregression tested associations between cohort characteristics and performance metrics. Results: This analysis included 1,595,977 women aged 20-75 years, enrolled in studies between 1976-2015, with 19,062 (1.2%) invasive BC cases ascertained within 5 years from exposure assessment. Age-adjusted AUCs were similar across models and cohorts (pooled AUCs by model: 0.57-0.58), while E/O ratios varied substantially (pooled E/O ratios by model: 0.83-1.25). Overestimation was common among predicted high-risk individuals (>3%). No appreciable differences in model performance by cohort age, birth year, race, and variable missingness emerged. Calibration improved after assigning race-specific incidence rates. Conclusion: Existing BC risk prediction models provided similar risk discrimination across multiple cohorts, although there was overestimation of risk for high-risk individuals. Performance variation across cohorts was not driven by specific characteristics, which supports development of a unified risk model for diverse populations that leverages appropriate incidence rates.

14
Integrated Clinicogenomic Risk Modeling for Metachronous Second Primary Cancers

Amsalem, J.; Ostrovnaya, I.; Marderstein, A. R.; Liu, Y. L.; Perea-Chamblee, T.; Ravichandran, V.; Jee, J.; Conry, M.; Khurram, A.; Kemel, Y.; Kim, E.; Mukherjee, S.; Latham, A.; Banaszak, L.; Kundra, R.; Magunta, S.; Fong, C.; Buas, M. F.; Bandlamudi, C.; Bernstein, J.; Seshan, V.; Salles, G.; Mandelker, D.; Berger, M. F.; Solit, D. B.; Stadler, Z. K.; Carrot-Zhang, J.; Schultz, N.; Offit, K.; Joseph, V.

2026-06-22 genetic and genomic medicine 10.64898/2026.06.12.26355388 medRxiv
Top 0.1%
4.9%
Show abstract

Improvements in cancer survival have increased the burden of subsequent primary malignancies. We developed and validated a programmatic classifier of multiple primary cancers (MPC) to derive second cancer phenotypes at scale. Among 81,175 cancer patients, we identified 56 first-second cancer pairs, 22 of which exceeded SEER primary cancer incidence rates. Even after accounting for various known risk factors, substantial elevated risk persisted, even in established hereditary cancer pairs (breast-ovary, breast-pancreas, prostate-pancreas), suggesting that current screening protocols do not adequately account for MPC susceptibility. To address this limitation, we built machine-learning models integrating rare germline variants, polygenic risk scores, treatment exposures, and demographic features to predict site-specific second primaries in breast and prostate cancer survivors. These models accurately predicted second ovarian and pancreatic cancers across a long follow-up period (15-year time-dependent AUC 0.70). This is the first systematic, pan-cancer integration of clinicogenomic factors for early prediction of second-primary malignancies. Our framework enables individualized risk estimation, enhanced targeted surveillance, and cancer prevention amongst a growing population of cancer survivors.

15
Inequalities in Colorectal Cancer Screening: Combining MAIHDA with Difference-in-Differences to Assess Programme Effects Across Population Subgroups

Jolidon, V.; Delaruelle, K.; Kawachi, I.; Cullati, S.; Bell, A.; Holman, D.

2026-07-18 health policy 10.64898/2026.07.16.26358242 medRxiv
Top 0.1%
4.8%
Show abstract

Background: Research consistently shows that colorectal cancer (CRC) screening uptake is socially patterned; however, sociodemographic determinants are usually analysed separately, overlooking how multiple social conditions jointly shape inequalities. This also applies to policy research, where heterogeneity in screening programme effects remains underexplored. Methods: Using data from the European Health Interview Survey (2014 and 2019; n=201,214; 24 countries), we applied Multilevel Analysis of Individual Heterogeneity and Discriminatory Accuracy (MAIHDA) to analyse CRC screening uptake across 72 subgroups defined by sex, education, living arrangement and employment. To assess heterogeneity in screening programme effects, we combined MAIHDA with difference-in-differences (MAIHDA-DiD). Results: MAIHDA revealed inequalities in uptake: lower- and middle-educated men, whether employed or unemployed, had the lowest uptake, whereas men and women not living alone, retired or living with disability, had the highest uptake. Lower-educated homemaker women were the only female group with below-average uptake. MAIHDA-DiD showed that programmes increased overall uptake but did not produce larger gains among groups with lower pre-intervention uptake, and therefore did not reduce inequalities. Instead, programmes generated above-average increases among groups with higher pre-intervention uptake, particularly lower- and middle-educated men and women not living alone and retired. Living arrangement explained more variation in programme effects than other factors, with individuals living alone benefiting less from the programmes. Conclusion: CRC programmes did not reduce (and may have widened) inequalities, underscoring the need for equity-focused strategies in population-based screening. By extending MAIHDA with difference-in-differences, this study introduces a novel approach for evaluating heterogeneous policy effects in public health.

16
APOBEC3G expression marks a TMB-high, T cell-inflamed tumor state and is associated with response to immune checkpoint blockade in multiple cancer cohorts

Butler, K.; Yesudhas, D.; Lone, B.; Banday, A. R.

2026-07-13 genomics 10.64898/2026.07.08.737369 medRxiv
Top 0.1%
4.7%
Show abstract

Immune checkpoint therapies have transformed clinical practice; however, reliable biomarkers to predict response remain limited. Tumor mutational burden (TMB) has emerged as an important biomarker because it is thought to reflect neoantigen load, yet its predictive utility has been inconsistent. This limitation may partly arise because TMB primarily captures tumor-intrinsic immunogenicity, which is heterogeneous and does not fully reflect the state of antitumor immunity. To identify transcriptomic surrogates that capture both high mutational burden and antitumor immune activation, we investigated whether mRNA expression of mutagenic APOBEC3 family members could serve as surrogates for high TMB and T cell-rich tumors. Using a pan-cancer computational framework, we evaluated the association of four APOBEC3 genes with mutational burden, neoantigen load, immune infiltration, and immune checkpoint blockade response. Among APOBEC3A, APOBEC3B, APOBEC3G, and APOBEC3H, APOBEC3G emerged as the strongest and most consistent marker of a TMBhighCD8high and NeoantigenhighCD8high tumor phenotypes. Single-cell analyses further demonstrated that APOBEC3G is enriched in both malignant cells and T cells compared with other APOBEC3 family members, with APOBEC3G-positive CD8+ T cells exhibiting elevated activation markers including GZMB and IFNG. Importantly, retrospective analyses of 50 immune checkpoint blockade cohorts showed that APOBEC3G had the most consistent association among APOBEC3 family members with treatment response and clinical outcomes. Together, these findings identify APOBEC3G as a candidate transcriptomic marker of a TMB-associated, T cell-inflamed tumor state linked to immune-checkpoint blockade benefit, warranting further prospective validation.

17
A germline KDM3C polymorphism impairs DNA repair and sensitizes to chemoradiotherapy

Hasan, A.; Demidova, E. V.; Priyadarshini, P.; Czyzewicz, P.; Gathuka, L.; Murayama, T.; Zhou, Y.; Kiss, Z. A.; Shastry, R. K.; Andrake, M.; Hearne, G.; Devarajan, K.; Wu, C.; Shah, A.; Schultz, B. M.; Connolly, D. C.; Rosen, G. L.; Canadas, I.; Liu, J. C.; Burtness, B. A.; Smith, J. J.; Dunbrack, R. L.; Golemis, E. A.; Whetstine, J. R.; Meyer, J. E.; Arora, S.

2026-08-31 genetic and genomic medicine 10.64898/2026.08.26.26360896 medRxiv
Top 0.1%
4.4%
Show abstract

Chemoradiotherapy (CRT) is the standard-of-care therapy for many solid malignancies, yet predictive biomarkers of treatment response remain limited. We identified a germline single nucleotide polymorphism (SNP) in an intrinsically disordered region of the lysine demethylase KDM3C/JMJD1C (p.S464T) that is associated with CRT outcomes in locally advanced rectal cancers (LARC) and head and neck squamous cell carcinoma (LA-HNSCC). In silico modeling with AlphaFold predicted S464T substitution influenced interaction between phosphorylated KDM3C and RNF8 FHA domain. In cellular models, conversion of S464 to T464 increased sensitivity to DNA-damaging agents. S464T substitution impaired damage-induced MDC1-RAP80 signaling and downstream RAP80-BRCA1 colocalization. SNP carrying cells impaired DNA repair causing genotoxic stress that is associated with increased cGAS-cGAMP innate immune signaling and increased apoptosis. Population analyses with the SNP highlighted an increase incidence of UV-induced skin and other cancers, linking inherited variation in the chromatin regulatory gene KDM3C to genome instability, cancer risk, and therapeutic vulnerability.

18
Cardiovascular Risk in BRCA1/2 Mutation Carriers: A Matched Cohort Study of Breast Cancer Survivors

Dehghan Manshadi, M.; Manouchehri, N.; Hubbert, L.; Liljegren, A.; Manouchehrinia, A.; Linder-stragliotto, C.; Rantala, J.; Hedayati, E.; Kiani, N.

2026-07-27 epidemiology 10.64898/2026.07.26.26350298 medRxiv
Top 0.1%
4.4%
Show abstract

Introduction: Cardiovascular disease (CVD) is a leading non-cancer cause of morbidity among breast cancer (BC) survivors. Among them, women carrying germline BRCA1 or BRCA2 mutations (BRCA-BC) may be at particular risk of CVD, but evidence is inconsistent. The objective of this study is to determine whether BRCA-BC independently influences the risk for CVD after BC diagnosis in the Stockholm-Gotland region in Sweden (2008-2019). Methods: In this registry-based cohort study, we used exact matching on age at diagnosis, tumor stage, laterality, and pre-existing CVD or risk factors to construct 32 matched (1:1) subgroups. Multi-state Cox proportional hazards models estimated hazard ratios (HRs) for transitions from BC diagnosis to first cardiovascular event, while accounting for competing risks of distant metastasis or non-cardiovascular death. Results: In matched subgroups, BRCA-BC experienced fewer CVD (6.4% vs. 11.2% (IQR 9.4%-12.2%), but significantly more competing events (22.3% vs. 10.1% (IQR 8.9%-11.3%); p<0.05. Multi-state Cox models revealed an inverse association between BRCA-BC status and the first cardiovascular event (HR<1), but a higher hazard of the competing risk. Cardiovascular events clustered in the first year after BC diagnosis, especially among BRCA-BC, suggesting truncated time at risk. Conclusion: BRCA-BC did not demonstrate increased cardiovascular risk after BC diagnosis. The apparent inverse association with CVD likely reflects the high incidence of competing risks, which limit the window for CVD to manifest. A small subgroup of long-term BRCA-BC survivors may represent biologically distinct individuals with different cardiovascular susceptibility, warranting further investigation.

19
Distinct Evolutionary Trajectories in Early Lung Adenocarcinoma: Age Related Pathway with Epidermal Growth Factor Receptor with Genome Doubling and Smoking-Driven Pathway

Goto, A.; Nakaoka, H.; Yoshida, M.; Koyama, K.; Miyabe, k.; Zhou, J.; Umakoshi, M.; Takashima, S.; Imai, K.; Minamiya, Y.; Nishikawa, K.; Matsubara, D.; Inoue, I.; Sugimura, H.; Ishikawa, Y.

2026-07-23 pathology 10.64898/2026.07.21.26358543 medRxiv
Top 0.1%
4.3%
Show abstract

Lung adenocarcinoma a progress from preinvasive lesions to invasive cancer; however, early evolutionary events in adenocarcinoma in situ (AIS) and minimally invasive adenocarcinoma (MIA) remain poorly defined, particularly in East Asian populations enriched for EGFR mutations. We performed whole-exome sequencing on 67 Japanese patients (38 AIS, 29 MIA), including multiregion sampling in 13 cases to identify EGFR as predominant driver (61.2%), followed by RBM10 (19.4%) and TP53 (9.0%). Two evolutionary trajectories emerged: age-related and smoking-driven pathways. In the former, EGFR-mutant tumors frequently exhibited early whole-genome doubling (WGD) (24.4%) with clock-like signature. The smoking-driven pathway, typically involving KRAS mutations, displayed a tobacco-associated signature. Multiregion sequencing revealed that driver mutations (EGFR, KRAS, and MET) were shared trunk events across in situ and invasive regions, while secondary alterations arose subclonally. This study defines the genomic evolution of early lung adenocarcinoma in Japanese patients, identifying two evolutionary trajectories: an age-related pathway with EGFR-linked genome doubling and a smoking-driven pathway involving KRAS mutations. These findings elucidate mechanisms underlying progression from preinvasive lesions to invasive cancer in Asian populations.

20
The Effect of Marital Status on Suicide Risk Among Patients with Breast Cancer: A Population-Based sIPTW Competing Risk Analysis

Zou, X.; Shi, J.

2026-07-04 oncology 10.64898/2026.07.01.26357044 medRxiv
Top 0.1%
4.3%
Show abstract

Background: Breast cancer survivors often experience psychological distress that may increase suicide risk. Marital status, a proxy for social support, may influence this risk, but its role within a competing-risk framework is unclear. This study examined the association between marital status and suicide mortality and assessed modification by socioeconomic and geographic factors. Methods: This is a population-based cohort study using SEER data, including adults diagnosed with primary breast cancer from 2000 to 2022. Marital status was classified as married/partnered or unmarried/non-partnered. Baseline characteristics were balanced using subdistribution inverse probability of treatment weighting (sIPTW). Suicide mortality was analyzed using sIPTW-weighted Fine-Gray competing-risk models, treating non-suicide deaths as competing events. Landmark, subgroup, interaction, and sensitivity analyses were performed. Results: Among 825,047 patients, 40.7% were unmarried. Covariates were well balanced after weighting (SMD <0.01). During follow-up, 529 suicide deaths occurred. Unmarried status was associated with higher suicide mortality (sHR = 1.34, 95% CI: 1.12-1.60). Male sex and estrogen receptor-negative tumors increased risk, while older age and non-White race were protective. Findings were consistent in Cox models (HR = 1.45) and sensitivity analyses (sHR = 1.42). Landmark analyses showed persistent associations at 1, 3, and 5 years. The association was attenuated in the highest income quartile but not modified by rural-urban status. Conclusions: Unmarried breast cancer patients had higher suicide mortality. These findings support integrating psychosocial assessment and targeted suicide prevention into survivorship care, especially for socially vulnerable groups.